[refactor] Refactoring AscendFusedMoE #1229

zzzzwwjj · 2025-06-15T10:13:26Z

What this PR does / why we need it?

This PR is used for resolved issue 1147

Move fused_moe code into one file fused_moe.py.
Integrate branch conditions into function get_fused_moe_state.

Does this PR introduce any user-facing change?

This PR has removed the env VLLM_ENABLE_MC2, because I think this env is useless, we can make judgments based on the current scenario without this env, it will only increase complexity.
This PR has removed the env USING_LCCL_COM, because this env has already expired.
additional_config.expert_tensor_parallel_size has already expired, and now we also use parameter enable_expert_parallel, consistent with the vLLM.

How was this patch tested?

wangxiyuan · 2025-06-16T01:05:10Z

Can you address the review comment from #1169

realliujiaxu · 2025-06-16T06:56:16Z

vllm_ascend/worker/model_runner_v1.py

why graph batch size starts from 4? this should be added into torchair config

This function is used to generate graph batch size automatically, if you want to control graph batch size, you can use torchair_graph_config.graph_batch_sizes to input it.

realliujiaxu · 2025-06-16T07:09:28Z

vllm_ascend/ops/fused_moe.py

this slice should be controlled by if not self.torchair_graph_enabled or is_prefill:

It is unnecessary to do this.

Signed-off-by: zzzzwwjj <1183291235@qq.com>

Yikun · 2025-06-17T06:18:51Z

@ganyi1996ppo @jianzs @ApsarasX Please also take a look

ganyi1996ppo · 2025-06-17T08:51:35Z

vllm_ascend/worker/model_runner_v1.py

Dose this means, we run all the scenario with full tokens?

indeed, if you don't set either of graph_batch_sizes_init and graph_batch_sizes.

Since we already support unbalanced mc2, why not just keep the bs1 graph here to miner the pressure of communication?

ganyi1996ppo · 2025-06-17T08:54:49Z

vllm_ascend/models/deepseek_dbo.py

Why there are only delete in dbo, have you verified the functionality of this?

ganyi1996ppo · 2025-06-17T08:59:41Z

vllm_ascend/models/deepseek_v2.py

Can you add comments here to illustrate the expected input and output shape of self.expert?

if you pass in shared_experts, self.expert will return a tuple: (router_hidden_states, shared_hidden_states), else will return a tensor of router_hidden_states.

I will add this note in code at next PR.

ganyi1996ppo · 2025-06-17T09:48:34Z

The dbo path seems have some issue in this PR, just discussed with @zzzzwwjj , we quick merge this code first for further release procedure, the fix PR will be filed asap after this PR merged.

ganyi1996ppo · 2025-06-17T09:53:13Z

@zzzzwwjj Please cherry-pick this PR to the dev branch

This PR is used for resolved [issue 1147](vllm-project#1147) 1. Move fused_moe code into one file `fused_moe.py`. 2. Integrate branch conditions into function `get_fused_moe_state`.  1. This PR has removed the env `VLLM_ENABLE_MC2`, because I think this env is useless, we can make judgments based on the current scenario without this env, it will only increase complexity. 2. This PR has removed the env `USING_LCCL_COM`, because this env has already expired. 3. `additional_config.expert_tensor_parallel_size` has already expired, and now we also use parameter `enable_expert_parallel`, consistent with the vLLM.   Signed-off-by: zzzzwwjj <1183291235@qq.com>

This PR is used for resolved [issue 1147](vllm-project#1147) 1. Move fused_moe code into one file `fused_moe.py`. 2. Integrate branch conditions into function `get_fused_moe_state`.  1. This PR has removed the env `VLLM_ENABLE_MC2`, because I think this env is useless, we can make judgments based on the current scenario without this env, it will only increase complexity. 2. This PR has removed the env `USING_LCCL_COM`, because this env has already expired. 3. `additional_config.expert_tensor_parallel_size` has already expired, and now we also use parameter `enable_expert_parallel`, consistent with the vLLM.   Signed-off-by: zzzzwwjj <1183291235@qq.com> Signed-off-by: ganyi <pleaplusone.gy@gmail.com>

### What this PR does / why we need it? This PR is the cherry-pick of the PR #1229 which have already merged into the main branch. This PR is used for resolved [issue 1147](#1147) 1. Move fused_moe code into one file `fused_moe.py`. 2. Integrate branch conditions into function `get_fused_moe_state`. ### Does this PR introduce _any_ user-facing change? 1. This PR has removed the env `VLLM_ENABLE_MC2`, because I think this env is useless, we can make judgments based on the current scenario without this env, it will only increase complexity. 2. This PR has removed the env `USING_LCCL_COM`, because this env has already expired. 3. `additional_config.expert_tensor_parallel_size` has already expired, and now we also use parameter `enable_expert_parallel`, consistent with the vLLM. ### How was this patch tested? CI passed Signed-off-by: zzzzwwjj <1183291235@qq.com> Signed-off-by: ganyi <pleaplusone.gy@gmail.com> Co-authored-by: zzzzwwjj <34335947+zzzzwwjj@users.noreply.github.com>

### What this PR does / why we need it? Add `max_num_tokens_across_dp` to AscendMetadata to fix dp This pr fixes the bug introduced by #1229, which add an arg `max_num_tokens_across_dp` when dp_size > 1. Signed-off-by: MengqingCao <cmq0113@163.com>

### What this PR does / why we need it? Add `max_num_tokens_across_dp` to AscendMetadata to fix dp This pr fixes the bug introduced by vllm-project#1229, which add an arg `max_num_tokens_across_dp` when dp_size > 1. Signed-off-by: MengqingCao <cmq0113@163.com>

### What this PR does / why we need it? This PR is used for resolved [issue 1147](vllm-project#1147) 1. Move fused_moe code into one file `fused_moe.py`. 2. Integrate branch conditions into function `get_fused_moe_state`.  ### Does this PR introduce _any_ user-facing change? 1. This PR has removed the env `VLLM_ENABLE_MC2`, because I think this env is useless, we can make judgments based on the current scenario without this env, it will only increase complexity. 2. This PR has removed the env `USING_LCCL_COM`, because this env has already expired. 3. `additional_config.expert_tensor_parallel_size` has already expired, and now we also use parameter `enable_expert_parallel`, consistent with the vLLM.  ### How was this patch tested?  Signed-off-by: zzzzwwjj <1183291235@qq.com>

### What this PR does / why we need it? Add `max_num_tokens_across_dp` to AscendMetadata to fix dp This pr fixes the bug introduced by vllm-project#1229, which add an arg `max_num_tokens_across_dp` when dp_size > 1. Signed-off-by: MengqingCao <cmq0113@163.com>

### What this PR does / why we need it? This PR is used for resolved [issue 1147](vllm-project#1147) 1. Move fused_moe code into one file `fused_moe.py`. 2. Integrate branch conditions into function `get_fused_moe_state`.  ### Does this PR introduce _any_ user-facing change? 1. This PR has removed the env `VLLM_ENABLE_MC2`, because I think this env is useless, we can make judgments based on the current scenario without this env, it will only increase complexity. 2. This PR has removed the env `USING_LCCL_COM`, because this env has already expired. 3. `additional_config.expert_tensor_parallel_size` has already expired, and now we also use parameter `enable_expert_parallel`, consistent with the vLLM.  ### How was this patch tested?  Signed-off-by: zzzzwwjj <1183291235@qq.com>

### What this PR does / why we need it? Add `max_num_tokens_across_dp` to AscendMetadata to fix dp This pr fixes the bug introduced by vllm-project#1229, which add an arg `max_num_tokens_across_dp` when dp_size > 1. Signed-off-by: MengqingCao <cmq0113@163.com>

github-actions bot added module:ops module:core module:quantization labels Jun 15, 2025

zzzzwwjj force-pushed the main branch from a31973f to 6614299 Compare June 15, 2025 13:28

zzzzwwjj force-pushed the main branch 2 times, most recently from 92c8998 to 494088d Compare June 16, 2025 07:14

realliujiaxu reviewed Jun 16, 2025

View reviewed changes

[refactor] Refactoring AscendFusedMoE

fa5574c

Signed-off-by: zzzzwwjj <1183291235@qq.com>

zzzzwwjj force-pushed the main branch from 494088d to fa5574c Compare June 16, 2025 18:46

wangxiyuan mentioned this pull request Jun 17, 2025

support fused_moe_allgather_ep #1220

Closed

Yikun added the ready read for review label Jun 17, 2025

zzzzwwjj requested a review from realliujiaxu June 17, 2025 06:48

realliujiaxu approved these changes Jun 17, 2025

View reviewed changes

wangxiyuan approved these changes Jun 17, 2025

View reviewed changes

ganyi1996ppo reviewed Jun 17, 2025

View reviewed changes

vllm_ascend/models/deepseek_dbo.py Outdated

Copy link

Collaborator

ganyi1996ppo Jun 17, 2025

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why there are only delete in dbo, have you verified the functionality of this?

ganyi1996ppo reviewed Jun 17, 2025

View reviewed changes

ganyi1996ppo approved these changes Jun 17, 2025

View reviewed changes

ApsarasX approved these changes Jun 17, 2025

View reviewed changes

ganyi1996ppo merged commit 23ca68d into vllm-project:main Jun 17, 2025
24 checks passed

ganyi1996ppo mentioned this pull request Jun 17, 2025

[v0.9.1][refactor] Refactoring AscendFusedMoE (#1229) #1264

Merged

This was referenced Jun 18, 2025

[DP] Add max_num_tokens_across_dp to AscendMetadata to fix dp and update example #1273

Merged

[v0.9.1][DP] Tiny fix of dp and update example #1277

Closed

This was referenced Jun 18, 2025

Support Deepseek w4a8 quantization #1182

Closed

[Bug]: deepseek-R1-w8a8 VLLM_ENABLE_MC2=1 error #1243

Closed

sdmyzlp mentioned this pull request Jun 25, 2025

Br fix multi stream moe #1417

Closed

[refactor] Refactoring AscendFusedMoE #1229

[refactor] Refactoring AscendFusedMoE #1229

Uh oh!

Conversation

zzzzwwjj commented Jun 15, 2025

What this PR does / why we need it?

Does this PR introduce any user-facing change?

How was this patch tested?

Uh oh!

wangxiyuan commented Jun 16, 2025

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Yikun commented Jun 17, 2025

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

ganyi1996ppo commented Jun 17, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

Uh oh!

ganyi1996ppo commented Jun 17, 2025

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

6 participants

ganyi1996ppo commented Jun 17, 2025 •

edited

Loading